Papers with dataset construction
Efficient Online Scalar Annotation with Bounded Support (P18-1)
Copied to clipboard
| Challenge: | Existing methods for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments are not shown. |
| Approach: | They propose a method for efficiently eliciting scalar annotations by human judgments. |
| Outcome: | The proposed method leads to increased correlation with ground truth, suggesting it is an improved mechanism for dataset creation and manual system evaluation. |
RethinkingTMSC: An Empirical Study for Target-Oriented Multimodal Sentiment Classification (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have shown that current TMSC systems rely on textual information, and the progress in tackling this task has slowed down. |
| Approach: | They propose to integrate both visual and textual information to improve the performance of TMSC by considering multimodal information. |
| Outcome: | The proposed model integrates both visual and textual information to improve performance. |
BlendX: Complex Multi-Intent Detection with Blended Patterns (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets such as MixATIS and MixSNIPS have limitations in their formulation. |
| Approach: | They propose a set of multi-intent detection datasets that feature more diverse patterns than their predecessors. |
| Outcome: | The proposed datasets feature more diverse patterns than their predecessors and are more complex and diverse than existing datasets. |
GLoHBCD: A Naturalistic German Dataset for Language of Health Behaviour Change on Online Support Forums (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing motivational interviewing methods lack the deep understanding of user utterances that is essential to the spirit of motivational interviews. |
| Approach: | They propose to use a German dataset of naturalistic language around health behaviour change to examine the motivational state of the user. |
| Outcome: | The proposed dataset of naturalistic language around health behaviour change is based on a weight loss forum in germany and is evaluated using theoretically grounded motivational interviewing categories. |
Pula: Training Large Language Models for Setswana (2025.naacl-long)
Copied to clipboard
| Challenge: | Setswana is a Bantu language spoken by an estimated five to ten million people worldwide. |
| Approach: | They propose to make setswana-based models available for the first time using data available from setswa and setswegian databases. |
| Outcome: | The proposed models outperform GPT-4o and Gemini 1.5 Pro on English-Setswana translation tasks and achieve state-of-the-art performance on Setswanan reasoning tasks. |
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)
Copied to clipboard
| Challenge: | Almost all popular summarization datasets do not come with inherent quality assurance guarantees. |
| Approach: | They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data. |
| Outcome: | The proposed metrics can be inexpensive heuristics for detecting generically low quality examples. |
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh (2025.acl-long)
Copied to clipboard
| Challenge: | Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains. |
| Approach: | They propose to open-source a large-scale instruction-following dataset covering key institutional and cultural knowledge relevant to Kazakhstan. |
| Outcome: | The proposed dataset improves LLMs’ understanding of procedural, legal, and structural governance topics. |
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur’anic Tafsir (2025.emnlp-main)
Copied to clipboard
| Challenge: | An annotated Tafsir ontology and a collection of 15 structured Tafsian books are presented in this paper. |
| Approach: | They propose a framework for retrieval and question-answering Tafsir data that spans the entire pipeline from dataset construction through evaluation and error analysis. |
| Outcome: | The proposed framework achieves 69.52% accuracy and 74.36% correctness overall, though multi-hop and context-dependent questions remain challenging. |
Personality Understanding of Fictional Characters during Book Reading (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to predict characters' personalities have not been studied in the NLP field due to the lack of appropriate datasets mimicking the process of book reading. |
| Approach: | They propose a dataset to predict characters' personalities that uses an exhaustive vocabulary of personality traits as targets. |
| Outcome: | The proposed dataset is efficient and accurate and relies on long-term context to achieve accurate predictions for both machines and humans. |
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios (2025.emnlp-main)
Copied to clipboard
| Challenge: | a number of tools are used to perform complex tasks, but the tool utilization process can cause errors. |
| Approach: | They propose a critique evaluation benchmark for tool learning that analyzes function-calling errors on tool evaluation benchmarks. |
| Outcome: | The proposed critique evaluation benchmark holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. |
Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)
Copied to clipboard
| Challenge: | Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm. |
| Approach: | They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning. |
| Outcome: | The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning. |
Batayan: A Filipino NLP benchmark for evaluating Large Language Models (2025.acl-long)
Copied to clipboard
Jann Railey Montalan, Jimson Paulo Layacan, David Demitri Africa, Richell Isaiah S. Flores, Michael T. Lopez Ii, Theresa Denise Magsajo, Anjanette Cayabyab, William Chandra Tjhi
| Challenge: | Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. |
| Approach: | They propose a benchmark that systematically evaluates LLMs across three key natural language processing competencies: understanding, reasoning, and generation. |
| Outcome: | The proposed benchmark covers eight tasks covering Tagalog and code-switched Taglish utterances. |
ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning (2025.emnlp-main)
Copied to clipboard
Rui Wang, Bohao Li, Xiyang Dai, Jianwei Yang, Yi-Ling Chen, Zhen Xing, Yifan Yang, Dongdong Chen, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang
| Challenge: | Existing approaches to adapt image-focused models for video understanding have not been successful in analyzing long video sequences. |
| Approach: | They propose a video instruction dataset that outperforms existing video instruction data for fine-tuning MLLMs by incrementally increasing input context length. |
| Outcome: | The proposed model outperforms existing models on video benchmarks and outperformed proprietary models on VideoMME even with a compact 7B model. |